This second iteration of SigLIP 2 introduces SigLIP 2, a family of new multilingual vision-language encoders that build on the success of the original SigLIP, and extends the original image-text training objective with several prior, independently developed techniques into a unified recipe.
A new artificial intelligence model, DeepSeek-R1, is introduced, demonstrating that the reasoning abilities of large language models can be incentivized through pure reinforcement learning, removing the need for human-annotated demonstrations.
A new model based onpi 0.5 is described that uses co-training on heterogeneous tasks to enable broad generalization and is demonstrated for the first time that an end-to-end learning-enabled robotic system can perform long-horizon and dexterous manipulation skills, such as cleaning a kitchen or bedroom, in entirely new homes.
A novel post-training recipe significantly improves the math, chat, instruction-following and multilingual abilities, making Gemma3-4B-IT competitive with Gemma2-27B-IT and Gemma3-27B-IT comparable to Gemini-1.5-Pro across benchmarks.
This work introduces GR00T N1, an open foundation model for humanoid robots that outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments and deploys the model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks.
Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, it is observed that FLUX.1 Kontext improved preservation of objects and characters, leading to greater robustness in iterative workflows.
AlphaEvolve is an evolutionary coding agent that substantially enhances capabilities of state-of-the-art LLMs on highly challenging tasks such as tackling open scientific problems or optimizing critical pieces of computational infrastructure.
The Cosmos World Foundation Model Platform is presented to help developers build customized world models for their Physical AI setups and position a world foundation model as a general-purpose world model that can be fine-tuned into customized world models for downstream applications.
With an improved framework for model development and evaluation, a large language model is shown to provide answers to medical questions that are comparable or preferred with respect to those provided by human physicians.
This work demonstrates how self-supervised learning from web-scale data and a small amount of robot interaction data can yield a world model capable of planning in the physical world.
The SCARE 2025 guideline provides an up-to-date framework for surgical case reports in the era of AI and adds specific reporting criteria for AI to ensure that any use of artificial intelligence in a case report is clearly documented, explained and discussed including with respect to bias and ethics.
The model, called CUT3R (Continuous Updating Transformer for 3D Reconstruction), captures rich priors of real-world scenes: not only can it predict accurate pointmaps from image observations, but it can also infer unseen regions of the scene by probing at virtual, unobserved views.
This work builds the first Multi-Agent System Failure Taxonomy (MAST), a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks, and develops an LLM-as-a-Judge pipeline with high agreement with human annotations to enable scalable annotation.
Terminal-Bench 2.0 is presented: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows and shows that frontier models and agents score less than 65\% on the benchmark and conducts an error analysis to identify areas for model and agent improvement.
A guideline to transparently reporting the use of AI in any manuscript in general is presented and will evolve over time as technology, systems and behaviour evolve.
The development of PROBAST+AI is described, which may replace the original PROBAST tool and allows all key stakeholders to examine the quality, risk of bias, and applicability of any type of prediction model in the healthcare sector, irrespective of whether regression modelling or AI techniques are used.
The results show the significant potential of AI in personalizing learning, automating routine tasks, and providing access to knowledge, but also reveal serious risks of exacerbating social inequality and ethical dilemmas.
The STROCSS 2025 guideline provides an up-to-date framework for surgical observational studies in the era of AI and adds specific reporting criteria for AI to ensure that any use of artificial intelligence in a surgical observational study is clearly documented, explained and discussed including with respect to bias and ethics.
A novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge is proposed and evaluated in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning.
OpenVLA, a 7B-parameter open-source VLA trained on a diverse collection of 970k real-world robot demonstrations, is introduced and it is shown that it can effectively fine-tune OpenVLA for new settings, with especially strong generalization results in multi-task environments involving multiple objects and strong language grounding abilities.
This work improves existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales and presents a novel transformer-based architecture for text-to-image generation that uses separate weights for the two modalities and enables a bidirectional flow of information between image and text tokens.
The new AlphaFold model demonstrates substantially improved accuracy over many previous specialized tools: far greater accuracy for protein–ligand interactions compared with state-of-the-art docking tools, much higher accuracy for protein–nucleic acid interactions compared with nucleic-acid-specific predictors and substantially higher antibody–antigen prediction accuracy.
Gemma 2, a new addition to the Gemma family of lightweight, state-of-the-art open models, ranging in scale from 2 billion to 27 billion parameters, delivers the best performance for their size, and even offers competitive alternatives to models that are 2-3 times bigger.
Results show that TorchDynamo is able to capture graphs more robustly than prior approaches while adding minimal overhead, and TorchInductor is able to provide a 2.41× training geometric mean speedup on an NVIDIA A100 GPU across 180+ real-world models, which outperforms six other compilers.
FineWeb is introduced, a 15-trillion token dataset derived from 96 Common Crawl snapshots that produces better-performing LLMs than other open pretraining datasets and FineWeb-Edu, a 1.3-trillion token collection of educational text filtered from FineWeb.
This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models, and presents comprehensive evaluations of safety and responsibility aspects of the models, alongside a detailed description of model development.
Recent improvements to Job Dispatcher are overviews, including its brand new website and documentation, enhanced visualisations, improved job management, and a rising trend of user reliance on the service from low- and middle-income regions.
This work introduces CONtrastive learning from Captions for Histopathology (CONCH), a visual-language foundation model developed using diverse sources of histopathology images, biomedical text and, notably, over 1.17 million image–caption pairs through task-agnostic pretraining, and represents a substantial leap over concurrent visual-language pretrained systems for histopathology.
It is shown that Chronos models significantly outperform other methods on datasets that were part of the training corpus; and have comparable and occasionally superior zero-shot performance on new datasets, relative to methods that were trained specifically on them.
Two below-threshold surface code memories on Willow, a distance-7 code and a distance-5 code integrated with a real-time decoder, indicate device performance that, if scaled, could realize the operational requirements of large-scale fault-tolerant quantum algorithms.
Code development continues in line with the Galaxy Project roadmap, with improvements to job scheduling and the user interface, and general purpose graphical processing units (GPGPU) access for cutting-edge methods, and licensed tool support.
This work introduces Tulu 3, a family of fully-open state-of-the-art post-trained models, alongside its data, code, and training recipes, serving as a comprehensive guide for modern post-training techniques.
Sarathi-Serve introduces chunked-prefills which splits a prefill request into near equal sized chunks and creates stall-free schedules that adds new requests in a batch without pausing ongoing decodes, and uniform batches in Sarathi-Serve ameliorate the imbalance between iterations resulting in minimal pipeline bubbles.
The status of InterPro is reported on, detailing new developments in the database, associated web interface and software, including the increased integration of structures predicted by AlphaFold and the enhanced description of protein families using artificial intelligence.
A unified software package CoverM is presented, which calculates several coverage statistics for contigs and genomes in an ergonomic and flexible manner and avoids unnecessary I/O overhead by calculating coverage statistics from streamed read alignment results.
Genie is introduced, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos, which enables users to act in the generated environments on a frame-by-frame basis despite training without any ground-truth action labels or other domain-specific requirements typically found in the world model literature.
OLMo is built, a competitive, truly Open Language Model, to enable the scientific study of language models and it is hoped this release will empower the open research community and inspire a new wave of innovation.
The BigCode project, an open-scientific collaboration focused on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder2, a large model that significantly outperforms other models of comparable size and makes the model weights available under an OpenRAIL license.
PaliGemma is an open Vision-Language Model that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model that achieves strong performance on a wide variety of open-world tasks.
Virchow is presented, the largest foundation model for computational pathology to date, and it is demonstrated that a large foundation model enables pan-cancer detection, achieving 0.95 specimen-level area under the receiver operating characteristic curve across nine common and seven rare cancers.
Article Galaxy Pages is a free service from Research Solutions, a company that offers access to content in collaboration with publishing partners, online repositories and discovery services.